Papers with jailbreak attacks
Dagger Behind Smile: Fool LLMs with a Happy Ending Story (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have attracted significant attention from jailbreak attacks . existing manual designs are either easily detectable or require intricate interactions with LLMs. |
| Approach: | They propose a happy ending attack that wraps up a malicious request in a scenario template . |
| Outcome: | The proposed attack wraps up a malicious request in a scenario template involving a positive prompt formed mainly via a happy ending, fooling LLMs into jailbreaking either immediately or at a follow-up malicious request. |
Virtual Context Enhancing Jailbreak Attacks with Special Token Injection (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing jailbreak attacks target the two phases of user interaction: prompt input and model computation. |
| Approach: | They propose a new tool that leverages special tokens to improve jailbreak attacks . they found that the tool can increase success rates of existing jailbreak methods by 40% . |
| Outcome: | The proposed solution can improve success rates of four widely used jailbreak methods by approximately 40% across various LLMs. |
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming (2025.acl-long)
Copied to clipboard
| Challenge: | Existing jailbreak techniques rely on single-round interactions, pro-Corresponding author. |
| Approach: | They propose a multi-turn safety alignment framework to address the challenge of securing large language models in multi-round interactions. |
| Outcome: | The proposed framework exhibits state-of-the-art attack capabilities while improving safety performance on safety benchmarks. |